On Pairwise Clustering with Side Information

نویسندگان

  • Stephen Pasteris
  • Fabio Vitale
  • Claudio Gentile
  • Mark Herbster
چکیده

Pairwise clustering, in general, partitions a set of items via a known similarityfunction. In our treatment, clustering is modeled as a transductive predictionproblem. Thus rather than beginning with a known similarity function, the functioninstead is hidden and the learner only receives a random sample consisting of asubset of the pairwise similarities. An additional set of pairwise side-informationmay be given to the learner, which then determines the inductive bias of ouralgorithms. We measure performance not based on the recovery of the hiddensimilarity function, but instead on how well we classify each item. We give tightbounds on the number of misclassifications. We provide two algorithms. The firstalgorithm SACA is a simple agglomerative clustering algorithm which runs in nearlinear time, and which serves as a baseline for our analyses. Whereas the secondalgorithm, RGCA, enables the incorporation of side-information which may lead toimproved bounds at the cost of a longer running time.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Online Active Constraint Selection For Semi-Supervised Clustering

Due to strong demand for the ability to enforce top-down structure on clustering results, semi-supervised clustering methods using pairwise constraints as side information have received increasing attention in recent years. However, most current methods are passive in the sense that the side information is provided beforehand and selected randomly. This may lead to the use of constraints that a...

متن کامل

Semi-Supervised Dimensionality Reduction Using Pairwise Equivalence Constraints

To deal with the problem of insufficient labeled data, usually side information – given in the form of pairwise equivalence constraints between points – is used to discover groups within data. However, existing methods using side information typically fail in cases with high-dimensional spaces. In this paper, we address the problem of learning from side information for high-dimensional data. To...

متن کامل

Active Learning of constraints using incremental approach in semi-supervised clustering

Semi-supervised clustering aims to improve clustering performance by considering user-provided side information in the form of pairwise constraints. We study the active learning problem of selecting must-link and cannot-link pairwise constraints for semi-supervised clustering. We consider active learning in an iterative framework; each iteration queries are selected based on the current cluster...

متن کامل

Spectral Active Clustering via Purification of the k-Nearest Neighbor Graph

Spectral clustering is widely used in data mining, machine learning and pattern recognition. There have been some recent developments in adding pairwise constraints as side information to enforce top-down structure into the clustering results. However, most of these algorithms are “passive” in the sense that the side information is provided beforehand. In this paper, we present a spectral activ...

متن کامل

Semi-supervised Clustering by Input Pattern Assisted Pairwise Similarity Matrix Completion

Many semi-supervised clustering algorithms have been proposed to improve the clustering accuracy by effectively exploring the available side information that is usually in the form of pairwise constraints. However, there are two main shortcomings of the existing semi-supervised clustering algorithms. First, they have to deal with non-convex optimization problems, leading to clustering results t...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:
  • CoRR

دوره abs/1706.06474  شماره 

صفحات  -

تاریخ انتشار 2017